Papers by Marie-Catherine de Marneffe
Label and Explanation Variation in LLM-Based Annotation: a Case Study in Natural Language Inference (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown considerable promise for annotation purposes, but questions remain about their ability to capture human label variation (HLV) label variation is genuine disagreement between annotators observed across NLP tasks. |
| Approach: | They investigate how label and explanation variation manifests within and across LLMs with respect to the Natural Language Inference task. |
| Outcome: | The proposed models generate label distributions similar to humans but exhibit distinct, idiosyncratic judgments and disagreement patterns. |
Identifying inherent disagreement in natural language inference (2021.naacl-main)
Copied to clipboard
| Challenge: | Natural language inference is the task of determining whether text is entailed, contradicted or unrelated to another piece of text. |
| Approach: | They propose to tease systematic inferences from disagreement items by capturing modes in annotations to simulate uncertainty in the annotation process. |
| Outcome: | The proposed approach performs statistically better than baselines on the CommitmentBank corpus in English. |
Do You Know That Florence Is Packed with Visitors? Evaluating State-of-the-art Models of Speaker Commitment (P19-1)
Copied to clipboard
| Challenge: | Existing models for speaker commitment fail to generalize to diverse linguistic constructions, highlighting directions for improvement. |
| Approach: | They evaluate two state-of-the-art speaker commitment models on the CommitmentBank . they analyze linguistic correlates of model error on a naturalistic dataset . |
| Outcome: | The proposed models perform well on some classes but fail to generalize to diverse linguistic constructions. |
Investigating Reasons for Disagreement in Natural Language Inference (2022.tacl-1)
Copied to clipboard
| Challenge: | Several disagreements in natural language inference (NLI) annotation are due to uncertainty in the sentence meaning, others to annotator biases and task artifacts. |
| Approach: | They propose a 4-way classification approach and a multilabel classification approach for detecting disagreements in natural language inference annotations. |
| Outcome: | The proposed model is more expressive and gives better recall of possible interpretations in the data. |
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection (2020.lrec-1)
Copied to clipboard
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, Daniel Zeman
| Challenge: | Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages. |
| Approach: | They describe version 2 of the universal guidelines and discuss major changes from UD v1 to UD 2 . they propose a morphological layer, a syntactic layer and a word segmentation layer . |
| Outcome: | The proposed treebanks are available for 90 languages and have been updated to meet the needs of multilingual parsers and researchers. |
Practical, Efficient, and Customizable Active Learning for Named Entity Recognition in the Digital Humanities (N19-1)
Copied to clipboard
Alexander Erdmann, David Joseph Wrisley, Benjamin Allen, Christopher Brown, Sophie Cohen-Bodénès, Micha Elsner, Yukun Feng, Brian Joseph, Béatrice Joyeux-Prunel, Marie-Catherine de Marneffe
| Challenge: | Scholars in interdisciplinary fields like the Digital Humanities are increasingly interested in semantic annotation of specialized corpora. |
| Approach: | They propose an active learning solution for named entity recognition that maximizes a custom model’s improvement per additional unit of manual annotation. |
| Outcome: | The proposed model reduces required annotation by 20-60% and outperforms a competitive active learning baseline. |
He Thinks He Knows Better than the Doctors: BERT for Event Factuality Fails on Pragmatics (2021.tacl-1)
Copied to clipboard
| Challenge: | Existing models for factuality prediction are lacking for English . Traditionally, event factualism is triggered by fixed properties of lexical items . |
| Approach: | They propose a model that exploits common surface patterns that correlate with factuality labels. |
| Outcome: | The proposed model achieves the best performance on four factuality datasets. |
Evaluating BERT for natural language inference: A case study on the CommitmentBank (D19-1)
Copied to clipboard
| Challenge: | Natural language inference datasets can identify premise-hypothesis relationship without observing premise . recasting of the CommitmentBank for NLI creates hypotheses that stand in entailment/contradiction/neutral relationship with premise. |
| Approach: | They propose to recast the CommitmentBank for NLI to stand in certain relationships with the premise . hypotheses are complements of clause-embedding verbs in each premise, rethinking the CommittedBank . |
| Outcome: | The proposed model performs well on the CommitmentBank with 85% F1 . however, the model does not capture the full complexity of pragmatic reasoning, authors say . |
Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations (2026.findings-acl)
Copied to clipboard
| Challenge: | Natural Language Inference (NLI) datasets often exhibit label variation. |
| Approach: | They extend LiTEx taxonomy to two NLI datasets and jointly analyze label variation and label variation. |
| Outcome: | The proposed model combines explanations as a lens to analyze variation in NLI annotations and examine individual differences in reasoning. |
Ecologically Valid Explanations for Label Variation in NLI (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Human label variation exists in many natural language processing tasks, including NLI . |
| Approach: | They build an English dataset of 1,415 ecologically valid explanations for 122 MNLI items . they find that people can systematically vary on their interpretation . |
| Outcome: | The proposed dataset contains 1,415 ecologically valid explanations for 122 items . the results show that people can vary on interpretation and highlight differences . |
Contextualized Embeddings for Enriching Linguistic Analyses on Politeness (2020.coling-main)
Copied to clipboard
| Challenge: | Current word embeddings in natural language processing do capture context and thus can be leveraged to enrich linguistic analyses. |
| Approach: | They propose a model which leverages pre-trained BERT to cluster contextualized representations of a word based on context in which it appears and labels of items it occurs in. |
| Outcome: | The proposed model can detect interpretable, finer-grained context patterns associated with (im)polite language. |
LiTEx: A Linguistic Taxonomy of Explanations for Understanding Within-Label Variation in Natural Language Inference (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing evidence of human label variation in Natural Language Inference (NLI) however, within-label variation is an additional challenge. |
| Approach: | They propose a linguistically-informed taxonomy for categorizing free-text explanations in English that captures different reasoning strategies behind NLI explanations with a particular focus on within-label variation. |
| Outcome: | The proposed taxonomy can be used to classify explanations in English using a linguistically-informed taxonomies. |